Back

Behavior Research Methods

Springer Science and Business Media LLC

Preprints posted in the last 90 days, ranked by how well they match Behavior Research Methods's content profile, based on 30 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
The e-Music Box Roma: an open research tool for accessible joint music making

F. Abalde, S.; Bigand, F.; Orciari, L.; Lorini, C.; E. Keller, P.; Parmiggiano, A.; Crepaldi, M.; Novembre, G.

2026-07-08 neuroscience 10.64898/2026.07.02.736121 medRxiv
Top 0.1%
33.9%
Show abstract

Joint music making offers an ecologically powerful framework for investigating human social interaction and synchronization. Yet, experimental paradigms often rely on traditional instruments that limit accessibility, reproducibility, and experimental control. In parallel, the use of music for therapy and rehabilitation is expanding, motivating the development of digital musical instruments that can serve research, educational, and clinical purposes. Here, we introduce the e-Music Box Roma (eMB Roma), an open, reproducible digital musical instrument designed to study music making behavior regardless of musical training. The eMB Roma plays preregistered music with tempo controlled by hand rotary movements. Building on the original e-Music Box (Novembre et al., 2015), the eMB Roma retains its intuitive rotary hand control while introducing major innovations: a fully open and 3D-printable design, modular hardware with integrated slider and button controls, polyphonic output with multiple simultaneous instruments, and MIDI compatibility. Additionally, a dedicated graphical user interface allows real-time monitoring, experiment control, device synchronization (like neuroimaging or motion capture devices), and both solo and joint music-making paradigms. The eMB Roma provides a flexible and accessible platform for research contexts, allowing experimental control, reproducibility, and future extensions. Its open design and modularity make it suitable not only for research but also for therapeutic, rehabilitation, and educational applications, where it can support personalized interventions and quantitative assessment of motor performance.

2
NC4touch: An open-source networked touchscreen apparatus for rodent behavioral testing

Modara, G.; Lester, A. W.; Cooke, M.; Schwein, I.; Li, V.; Zhang, J.; Wong, A.; Cid, L.; Snyder, J. S.; Madhav, M.

2026-08-04 animal behavior and cognition 10.64898/2026.07.30.741621 medRxiv
Top 0.1%
23.2%
Show abstract

In recent years, rodent touchscreen-based paradigms have gained popularity for their flexible task design, automated data collection and improved standardization. These devices minimize experimenter intervention and contribute to more replicable and reliable research. However, the high costs and proprietary hardware / software associated with commercially available systems present a financial barrier for smaller labs or researchers with limited resources. To address this issue, we developed "NC4Touch", an open-source scalable, modular rodent testing apparatus. It features three independent touchscreens and a feeding port for easy use in a wide array of tests comparable to those available with standard operant chambers, with the added benefit that visual stimuli can be highly customized. The system includes a user-friendly graphical interface that offers real-time control, task customization, video recording, and data management. We demonstrate its effectiveness in a visual discrimination task using both rats and mice. Several devices can be operated in parallel from the same computer and user interface, allowing high- throughput data collection. We provide detailed assembly instructions for the hardware and well- documented and easily-configurable software. We hope that the affordable open-source nature of NC4Touch will expand the scope of behavioral testing, allowing researchers to overcome traditional financial barriers and create a more collaborative community. Significance StatementTouchscreen-based testing is a powerful tool for assessing cognition in rodent models, however commercial systems remain cost-prohibitive for many researchers. NC4Touch is a rodent touchscreen apparatus that provides an open-source, customizable, and affordable alternative that enables high-throughput, cognitive testing in both mice and rats. By lowering financial barriers and promoting hackability, this system enables greater accessibility and collaboration in research.

3
Multisensory Continuous Psychophysics: Perceived Visual Object Location is Improved by Auditory Cues

Jörges, B.; Kim, J.-J.; Harris, L. R.

2026-06-08 animal behavior and cognition 10.64898/2026.06.03.729954 medRxiv
Top 0.1%
22.3%
Show abstract

Continuous Psychophysics, which couples a continuous stimulus with a continuous response, is a promising tool to break out of the confines of traditional designs based on discrete trials. In this pre-registered study, we explore to what extent this paradigm is useful in the study of multisensory integration. We expand on Tonelli et al.s (2025) seminal study by additionally examining the role of eye-movements, using a Kalman filter to estimate the sensory noise underlying behavioral tracking parameters and employing a virtual reality set-up. We immersed two cohorts of participants (n = 30 each) in a virtual meadow environment and asked them to continuously track a drone (Experiment 1) or a swarm of flies (Experiment 2) with a controller, while simultaneously recording their eye movements. We manipulated the reliability of visual cues using four levels of fog (from a completely clear view to impenetrable fog where no visual cues to the targets position were available) as well as the presence of sound cues emitted from the object (sound present/absent). The maximum correlation between stimulus and response was higher when sound was present in some conditions, particularly when visual uncertainty was high, while the tracking delay remained unaffected across all fog levels. Using a Kalman filter to estimate the underlying sensory noise, we found strong evidence that sensory noise was lower when sound was present than when sound was absent both for manual and for ocular tracking, particularly for those conditions with higher visual uncertainty. In exploratory analyses, we further show strong correlations between manual and ocular tracking in all measures (maximum correlation, tracking delay, sensory precision). However, when isolating the multisensory advantage, these correlations all but disappeared for maximum correlation and tracking delay, while remaining substantial for sensory precision. Similarly, behavioral tracking correlated generally strongly with underlying sensory noise, but much less so when it came to the advantage conferred by added sound cues. Our results show that continuous psychophysics is well-suited for the study of multisensory integration, particularly when a Kalman filter analysis is used to estimate sensory uncertainty from behavioral data.

4
PLFest: A Multi-Site Validation of an Open Platform for Visual and Cognitive Assessment

Penaloza, B.; Maniglia, M.; Munneke, J.; Green, C. S.; Seitz, A.

2026-06-28 neuroscience 10.64898/2026.06.22.733892 medRxiv
Top 0.1%
21.8%
Show abstract

Purpose: To evaluate the feasibility, validity, and scalability of PLFest, an open-source, Unity-based, cross-platform application designed for standardized, multi-site visual and cognitive assessment and training. Methods: Two hundred sixty participants (mean age = 23 years) were recruited across four university sites in the United States. Participants completed a battery of five visual assessments administered through PLFest, including visual acuity, contrast sensitivity, spatial frequency cutoff, contrast sensitivity at spatial-frequency cutoff, and visual search. Five cognitive assessments measuring visuospatial working memory, verbal working memory, fluid reasoning, inhibitory control, and selective attention were also administered. Descriptive statistics and performance distributions were examined and compared with normative data. Results: Visual acuity and contrast sensitivity measures closely matched previously reported normative values obtained using established clinical and psychophysical methods. Spatial frequency cutoff and visual search tasks produced stable threshold estimates while showing substantial inter-individual variability. Performance across all cognitive assessments was consistent with published validation studies of the corresponding tasks. Across the full battery, adaptive procedures demonstrated reliable convergence and generated well-distributed performance measures without evidence of substantial floor or ceiling effects. Importantly, these findings were observed across four geographically distributed testing sites using standardized consumer-grade tablet hardware. Conclusions: PLFest provides reliable and scalable assessment of visual and cognitive function using portable consumer devices. The platform supports standardized data collection across distributed research settings while maintaining performance characteristics consistent with established laboratory and clinical benchmarks. These findings support the use of PLFest as a reliable framework for large-scale studies of vision and cognition. Translational Relevance: By reducing dependence on specialized laboratory infrastructure and trained personnel, PLFest may facilitate broader access to visual and cognitive assessment, enabling large-scale research, screening, and future rehabilitation applications.

5
ResXR: validated infrastructure for reproducible studies of human behavior in Extended Reality

Bergstein, Y.; Barel, N.; Shai Basson, G.; Bromberg, O.; Schonberg, T.

2026-06-22 neuroscience 10.64898/2026.06.17.732869 medRxiv
Top 0.1%
15.3%
Show abstract

Extended Reality (XR) combines experimental control and ecological validity, yet behavioral XR research lacks shared infrastructure: building immersive experiments demands specialized engineering, and custom tools yield data in custom formats that other laboratories cannot readily reanalyze. We present ResXR (Research with XR), an open-source toolkit providing a path from immersive experiment to standardized dataset and quality report, running on standalone headsets. A Unity template records synchronized head, hand, eye, and face tracking with per-sample hardware timestamps; an independent Python pipeline validates quality, masks flagged intervals, and exports raw and derivative datasets in Motion-BIDS format with self-contained quality reports. Three ready-to-run paradigms span common behavioral designs. ResXR is an idea new to XR research: sensor data from consumer headsets must be empirically validated rather than taken from vendor documentation, grounding its schema and quality flags in stress-tested sensor behavior. Its aim is a transparent, community-extensible foundation for reproducible XR experimentation.

6
Initial Technical and Clinical Validation of Mobile Pupillometry with Virtual Reality: A Digital Biomarker for Screening Cognitive Function and Impairment

Brendler, A.; Fietz, J.; Bauer, A.; Pfahl, D.; Higgins, S.; Vidovic, E.; Brueckl, T.; BeCOME Working Group, ; Memory Clinic Working Group, ; Hupe, K.; Knop, M.; Spoormaker, V. I.

2026-07-17 neurology 10.64898/2026.07.15.26358187 medRxiv
Top 0.1%
12.5%
Show abstract

Cognitive impairment is a prevalent symptom extending from physiological ageing to disease. It commonly manifests itself in initial memory problems, progressing and co-occurring in more severe conditions such as Mild Cognitive Impairment, Alzheimer's Disease and Major Depressive Disorder. However, current non-invasive screening assessments either lack biological information or are invasive and restricted to specialized centers with complex and cost-intensive set-ups. Here, we conducted an initial validation of mobile pupillometry with Virtual Reality (VR) under experimental conditions as a digital biomarker for cognitive impairment by testing required biomarker-specific properties. For this purpose, we first assessed its construct validity by testing healthy participants (n=43) on an n-back task in VR while pupil size was measured. Mixed effects models revealed that similar to lab-based eye-tracking systems, pupil size increased in a sensible and distinguishable fashion as a function of working memory load. Second, to test the signal's reliability, the same participants were tested on the identical set-up two to three months after their first visit. We observed that the pupil response profile was highly stable over this period. Third, for its clinical validity, we examined patients (n=89) from three different cohorts with varying degrees of cognitive impairment and compared them to healthy control participants (n=81). Mixed-effects models indicated that pupil size was reduced as a function of cognitive impairment levels at higher cognitive load and that this effect was stronger pronounced with increasing age. In conclusion, we provide initial evidence for mobile pupillometry being a sensitive, reliable and clinically valid digital biomarker for cognitive functioning and impairment, which offers desirable properties due to its quick, automatized and location-independent set-up. Keywords: digital biomarker, mobile pupillometry, Virtual Reality, cognition, , Major Depressive Disorder, Mild Cognitive Impairment, Alzheimer's Disease

7
Robean: A Standalone Freeware for Automated Rodent Neurobehavioural Analysis with Integrated Tracking, Visualization, and Reporting

Mishra, V.; Verma, R.; Rajinikanth, P. S.; Kaundal, R. K.

2026-07-17 animal behavior and cognition 10.64898/2026.07.11.737999 medRxiv
Top 0.1%
8.1%
Show abstract

Quantitative analysis of rodent behaviour is fundamental to neuroscience, preclinical drug discovery and neurotoxicology research. Although several commercial and open-source software packages are available for behavioural assessment, some are expensive, some require programming expertise, and some provide limited flexibility for user-defined experimental configurations. To address these limitations, we developed Robean, a freely available standalone software platform for automated rodent neurobehavioural analysis from both live camera feeds and pre-recorded videos. Robean provides an intuitive graphical user interface that enables users to design experimental arenas, define custom analysis zones, perform spatial calibration, and automatically track rodent movement without requiring programming knowledge. The software currently supports automated analysis of three widely used behavioural paradigms: the Morris Water Maze, Elevated Plus Maze, and Open Field Test. Robean extracts behavioural metrics including escape latency, path efficiency, platform crossings, target quadrant preference, thigmotaxis, locomotor activity, zone occupancy, arm entries, and centre exploration specific to behavioural tests. In addition, the software generates trajectory maps, occupancy heatmaps, comma-separated value (CSV) datasets, comprehensive PDF reports, and batch study summaries for multiple experimental sessions. Developed using open-source software technologies and distributed as a standalone freeware application, Robean provides an accessible and reproducible solution for behavioural neuroscience laboratories. Its modular architecture facilitates future integration of additional behavioural paradigms and analytical modules, making it a flexible platform for automated rodent behavioural assessment.

8
When communication fails, physical effort increases but not to greater effect

Kadava, S.; Pouw, W.; Fuchs, S.; Holler, J.; Cwiek, A.

2026-07-10 physiology 10.64898/2026.07.06.736716 medRxiv
Top 0.1%
7.8%
Show abstract

Successful communication depends not just on what we say but also on how we repair it when understanding breaks down. People routinely respond to misunderstanding by gesturing more or speaking louder. Yet whether these responses actually reflect measurable bodily effort, and whether it helps communication, has not yet been directly tested. Extending the hypo- and hyper-articulation accounts to whole-body multimodal communication we ask: does regaining understanding require increased physical effort, and does that effort improve communicative outcomes? In a pre-registered biomechanical study, 61 pairs of participants conveyed concepts without conventionalized language, using gestures, vocalizations, or both, and repaired failed attempts. We measured communicative effort by capturing arm torque change (upper-limb effort), vocal amplitude (vocal effort), and postural adjustments (postural effort). Using Bayesian hierarchical models, we show that repair increases physical effort across all three measurements, but that performers communicate more over time rather than more forcefully at any particular moment in time. Upper-limb effort scale with the degree of misunderstanding but effort overall does not predict whether communication succeeds. We suggest that communicative effort on the level we measured here functions as an index of continued commitment rather than a mechanism for repairing misunderstanding.

9
Subtle Pupil-Size Changes Associated With Exploration Do Not Affect Visual Sensitivity

Claeys, W.; Ruuskanen, V.; Mathot, S.

2026-06-09 animal behavior and cognition 10.64898/2026.06.05.730459 medRxiv
Top 0.1%
7.7%
Show abstract

When we feel restless and easily distracted, continuously switching tasks (exploration), our pupils tend to be large. In contrast, when we are calmly focused on a single task (exploitation), our pupils tend to be small. According to the Adaptive Gain Theory (AGT), a switch from exploitation to exploration is associated with an increase in norepinephrine in the locus coeruleus, which in turn triggers pupil dilation. However, the AGT does not provide a functional explanation of why exploration triggers pupil dilation. One possibility is that visual sensitivity, which increases with pupil size, is especially important during exploration. We set out to provide evidence consistent with this functional explanation, as well as to replicate two key previous results. Participants performed a four-armed bandit task, which induces both exploration and exploitation behavior. During the task, participants also needed to detect an occasional and unpredictable near-threshold peripheral flash. We replicated two key results: pupils were larger during exploration than during exploitation; and increased pupil size (overall, independent of exploration status) was associated with increased visual sensitivity. However, most importantly, we did not find that visual sensitivity was higher during exploration than during exploitation; probably, the reliable-yet-tiny increase in pupil size during exploration was too small to affect visual sensitivity. We conclude that key previous results are replicable; however, common experimental paradigms, such as the four-armed bandit task, induce only small changes in exploration behavior. Therefore, more powerful paradigms are required in order to test functional explanations of pupil-size changes during exploration and exploitation.

10
Linking continuous behavior to aesthetic enjoyment in a walkable virtual-reality museum tour: effects of agency and a painting-level analysis framework

Sklyar, Y.; Hendler, S.; Schonberg, T.

2026-09-01 neuroscience 10.64898/2026.08.27.747505 medRxiv
Top 0.1%
7.4%
Show abstract

Museum visits typically follow curator-defined routes that constrain how visitors shape their own experience, yet choice is widely held to heighten engagement, autonomy, and enjoyment. Virtual reality (VR) offers a setting in which to study these processes because it combines ecological immersion with precise, continuous behavioral measurement. We investigated (i) whether VR- derived behavioral signals are associated with self-reported enjoyment during a virtual museum tour, and (ii) whether the level of agency afforded to visitors influences enjoyment. Forty-eight adults completed a room-scale, life-size VR tour (8 * 4 m) of seven paintings from the Tel Aviv Museum of Art, each accompanied by a synchronized audio guide. Synchronized gaze and head- position streams were logged continuously (50 Hz) and segmented into painting-level viewing episodes using a trial-and-tile pipeline that intersects each painting's trial interval with an empirically defined spatial window in front of the canvas. Participants were randomly assigned to one of three agency conditions, Active (choice before every artwork), Semi-Active (choice for the first three), or Passive (fixed route),while the artwork sequence was held identical. Self- reported enjoyment at the tour and painting levels did not differ reliably across agency conditions. Among VR-derived measures, gaze engagement during the audio guide showed the clearest (though modest) association with painting-level liking, whereas locomotion and pacing measures were weak and inconsistent predictors. Agency nonetheless reliably modulated several gaze- and time-based viewing measures. The findings reveal a dissociation between subjective enjoyment and the micro-structure of viewing, and establish a reusable framework for full-tour, painting-level behavioral analysis in immersive settings.

11
Computer Vision Scoring of Figure Copy and Recall

Woods, D. L.; Hall, K.; Jaramillo, I.; Blank, M.; Geraci, K.; Boghassian, A.; Pebler, P.

2026-06-11 neurology 10.64898/2026.06.10.26355298 medRxiv
Top 0.1%
7.3%
Show abstract

Objective. Figure copy and recall tests are sensitive measures of visuoconstruction and visual episodic memory, but their clinical is constrained by labor-intensive manual scoring. We developed and validated an automated, element-level scoring pipeline using Vertex AI object detection for the tablet-based figure copy and recall tasks in the California Cognitive Assessment Battery (CCAB). The automated scoring pipeline duplicated the scoring procedures used by expert manual raters. Methods. A normative sample of 2,011 community-dwelling adults aged 18-90 completed figure copy and delayed recall trials at baseline, with subsamples retested at 1 day and at 6, 18, and 30 months. Participants completed the drawings with their index finger on a tablet computer with finger position digitized to analyze the speed and timing of individual drawing strokes A convolutional object-detection model trained on the Vertex AI AutoML Vision platform identified each of twelve canonical figure elements in rendered drawings. Separate element presence and location scores were computed after homographically warping drawings onto a canonical template to produce trial-level Element, Location, and Total scores. To compare Vertex and human scores, Vertex AI and expert human raters independently scored 1500 randomly selected drawings to evaluate inter-rater agreement, including a common subset of 100 drawings scored by Vertex AI and all raters. Results. Total scores were virtually indistinguishable (r = 0.966) from human-human agreement (mean r = 0.971) as were Element presence scores (mean r = 0.959 vs. r = 0.963). Location-score agreement (r = 0.951) was slightly below the human-human mean (r = 0.972) due to pixel-level analysis by Vertex AI that was impossible for human raters. The Vertex pipeline showed no preferential advantage for the single expert rater who categorized Elements during training. Automated scores showed strong demographic gradients, age effects on Recall (r = -0.32) were approximately twice those in Copy conditions (r = -0.16). A Memory Cost score (Recall - Copy) showed a monotonic age-related decline from +0.40 z in the youngest subjects to -0.54 z in the oldest. Kinetic analysis revealed that drawing speed and efficiency showed significant age-related changes. Overnight test-retest reliability was high (Recall r = 0.72) and the Recall trial showed a large overnight learning effect ({Delta} = +1.18) that continued with repeated tests up to 30 months ({Delta} = +0.75).

12
Unraveling emotional signatures: comparing physiological methods and algorithm-based recognition of spontaneous emotional facial expressions

Kissler, J. M.; Scholz, S.

2026-08-10 neuroscience 10.64898/2026.08.05.742947 medRxiv
Top 0.1%
6.4%
Show abstract

Recognizing others emotions is central to social interaction. Traditional biological psychology infers emotional responding via laboratory measures, whereas contemporary computer vision algorithms claim to identify emotions unobtrusively from facial video. However, the validity of such algorithms for classifying spontaneous emotional responses occurring without explicit communicative intent remains debated. We compared established psychophysiological measures (EEG, facial EMG, EDA activity) with the open-source facial behavior toolkit OpenFace for classifying participants spontaneous responses during free viewing of happiness-inducing, disgust-inducing, and neutral pictures. Participants provided valence and arousal ratings and later selected the basic emotion that best matched their reaction which served as the classification criterion. Using within-participants single-trial support vector machine (SVM) classification, EEG achieved the highest accuracy (40%), followed by facial EMG (37%); OpenFace reached 36%. All methods except EDA exceeded chance performance (33.3%) and were lower compared to human raters (48%). Predictions declined slightly for across-participants SVMs, being at chance for OpenFace and EDA. The results indicate that in principle both, psychophysiological measures and video-derived facial action units, can capture diagnostically relevant aspects of emotional responding during picture viewing, but that their performance is limited when expressions are spontaneous and not produced for communicative purposes. Inter-individual variability in expressivity and physiological responding likely contributes to these limitations and should be considered when deploying automatic emotion recognition in research or applied settings.

13
Evidence that Dogs Can Use Temporal Difference in Odorant Arrival to Discriminate Odorant Mixtures

Downie, I.; Szyszka, P.; Hall, N. J.; Edwards, T. L.

2026-07-01 animal behavior and cognition 10.64898/2026.06.26.734637 medRxiv
Top 0.1%
5.5%
Show abstract

In turbulent environments, odorants from different sources arrive at different times, potentially providing cues for odor source segregation. In several invertebrate species, short differences in odorant onset enable freely moving animals to discriminate odorant mixtures. In vertebrates, however, studies of sensitivity to odorant onset asynchrony have been conducted under highly constrained sampling conditions, such as with odor delivery tightly coupled to respiration. In this study, we investigated whether domestic dogs could detect odorant onset asynchrony in odorant mixtures under conditions that preserve key features of natural odor sampling. Dogs performed a discrimination task in which odor stimuli were presented as ongoing pulse trains that began independently of animal behavior, avoiding artificial synchronization of odor delivery with sniff cycles. Dogs were trained to discriminate between mixtures of two odorants with synchronous onsets and mixtures with asynchronous onsets. Of the dogs trained, one was able to discriminate odorant onset asynchronies as short as 633 ms. Dogs also displayed sensitivity to auditory stimulus onset asynchrony, discriminating auditory asynchronies as short as 30 ms. These results provide the first demonstration of temporal sensitivity in canine olfaction and the first evidence that vertebrates can use odorant onset asynchrony under conditions that permit free odor sampling.

14
Effects of aging on multiple object tracking under normal and altered viewing conditions

Michaud, C.; Baures, R.; Soler, V.; Trotter, Y.; Vattier, V.; Rosito, M.; Peyrin, C.; Cottereau, B. R.

2026-06-30 animal behavior and cognition 10.64898/2026.06.25.734471 medRxiv
Top 0.1%
5.4%
Show abstract

Multiple object tracking (MOT) is a core function of dynamic visual attention that relies on the ability to simultaneously monitor several moving objects. Although MOT performance is known to decline with age, and to depend on efficient oculomotor strategies, how these processes interact across the adult lifespan and under degraded visual input remains poorly understood. Here, we examined the effects of aging on MOT under normal and gaze-contingent viewing conditions simulating central and peripheral visual field loss. Sixty participants aged 20-80 years completed a MOT task while eye movements were recorded, enabling characterization of performance and oculomotor behavior across five viewing conditions. Behavioral results revealed a continuous decline in tracking performance across adulthood, indicating a graded rather than categorical effect of age. Performance was strongly reduced by visual-field restrictions, with the largest impairments under central vision occlusion. Eye-tracking analyses showed that better performance was associated with greater reliance on centroid-based gaze strategies, consistent with distributed monitoring of target configurations. Critically, older adults relied more on focal, target-based tracking under conditions simulating peripheral vision loss, and less on centroid-based strategies; this shift was associated with poorer performance. In contrast, oculomotor behavior during full-field viewing was largely preserved across age. Together, these findings suggest that aging affects multiple object tracking through combined sensory, attentional, and oculomotor mechanisms. Beyond a reduction in capacity, age-related decline also reflects systematic changes in visual sampling strategies during dynamic tracking.

15
Cross-session generalization in automated behavioral tracking of Galleria mellonella larvae: comparison of classical computer vision, deep learning and generative domain adaptation

Smigielski, K.; Piorkowska, N.

2026-07-28 animal behavior and cognition 10.64898/2026.07.24.740557 medRxiv
Top 0.1%
5.4%
Show abstract

BackgroundAutomated behavioral tracking is increasingly used in biological and biomedical research; however, robustness across heterogeneous imaging conditions remains a major challenge. Domain shifts caused by changes in illumination, contrast, or acquisition setup can substantially degrade the performance of computer vision and deep learning models, limiting their practical applicability. This issue is particularly relevant for small-scale biological datasets, where extensive annotation and retraining are often impractical. MethodsWe developed a behavioral tracking framework for Galleria mellonella larvae combining a classical computer vision (CV) pipeline, a YOLOv8s-seg + ByteTrack deep learning pipeline, and generative domain adaptation methods. Behavioral recordings were collected in two independent experimental sessions under distinct illumination conditions (top and bottom lighting), creating a natural cross-session domain shift. The classical pipeline was based on contour detection, adaptive preprocessing, temporal smoothing, and trajectory reconstruction. Deep learning models were trained on 320 manually annotated frames and evaluated in both within-session and cross-session settings. To improve generalization, we investigated generative augmentation using StyleGAN2-ADA and unpaired image-to-image translation using CycleGAN. Tracking outputs were further analyzed through behavioral descriptors, including trajectories, traveled distance, velocity, spatial occupancy heatmaps, and directional movement patterns. ResultsThe classical CV pipeline achieved high detection performance in both recording sessions, with mean detection rates of 99.05% and 99.71%, respectively. A YOLOv8s-seg model demonstrated strong within-session performance but exhibited severe degradation under cross-session evaluation, with larval mask segmentation performance dropping to mAP@0.5 = 9.1%, confirming the presence of a substantial domain shift. Despite differences in detection methodology, behavioral metrics derived from YOLO and CV pipelines showed strong agreement at the group level (Pearson correlation r = 0.89; median distance ratio = 0.99). Generative augmentation with StyleGAN2-ADA did not yield meaningful gains in tracking robustness -- likely because baseline performance was already near ceiling -- whereas CycleGAN-based domain adaptation substantially reduced the domain gap and improved cross-session detection performance while preserving biologically relevant trajectory structures and spatial behavioral patterns. ConclusionsCross-session variability represents a critical challenge for automated behavioral tracking in biological experiments. Our results demonstrate that carefully designed classical computer vision approaches can achieve highly reliable tracking in small-data settings, while deep learning models require explicit strategies to address domain shift. Generative domain adaptation, particularly CycleGAN-based image translation, offers an effective solution for improving cross-session generalization without additional manual annotation. The proposed framework provides a robust foundation for scalable behavioral phenotyping of Galleria mellonella and other small biological model organisms.

16
What group averages conceal: functional heterogeneity in human eyeblink habituation

Perez, O. D.; Cancino, N.; Hermosilla, D.; Soto, F. A.; Vogel, E. H.

2026-06-25 animal behavior and cognition 10.64898/2026.06.21.733594 medRxiv
Top 0.1%
5.4%
Show abstract

In animal learning research, learning is often represented by plotting a behavioral measure as a function of training trials. A particularly clear case is habituation, a basic form of learning in which repeated presentation of a stimulus produces a decrement in responding. Although retention tests provide the strongest basis for evaluating durable habituation once short-lived performance effects have dissipated, the pattern of response change across stimulus repetitions, or habituation curve, remains theoretically and empirically relevant because it is used to characterize determinants of habituation, individual and clinical profiles, and functional forms, including linear, curvilinear, asymptotic, and mixed incremental-decremental patterns of responding. However, group averaged curves may conceal substantial individual heterogeneity. Here, we analyzed archived human eyeblink habituation data from 157 participants to ask whether the curve shape selected for the group average reflects the curve shapes observed at the individual level. Five candidate functions were fitted separately to each participant and to the corresponding group average. No single function characterized most individuals. More importantly, the model selected for the group average differed from the most frequent individual model in all four groups. When data were pooled across groups, the average favored a dual-process form, a shape that matched the individual plurality in none of them. Simulation analyses showed that averaging heterogeneous individual trajectories can itself produce a group curve that favors a more complex model. Our findings show that group averaged habituation curves should not be treated as direct descriptions of the typical individual trajectory.

17
Individual Differences in Color-Induced Visual Discomfort Reveal a Pupillary Dissociation Across Stimulus Dimensions

Meidan, R. Y.; Bonneh, Y. S.

2026-07-22 neuroscience 10.64898/2026.07.17.739245 medRxiv
Top 0.1%
5.3%
Show abstract

Visual discomfort (VD) is influenced by both spatial structure and chromatic context. Striped patterns are well-established triggers of discomfort and autonomic responses. In previous work, we showed that higher spatial frequencies and larger patterned areas elicit stronger pupillary constriction and greater discomfort, and that individuals with higher overall discomfort show shallower maximum constriction. The present study examined whether similar relationships appear when spatial structure is held constant, measured background luminance is kept within a narrow range, and the chromatic background varies. Participants viewed black horizontal stripes on 12 near-isoluminant colored backgrounds. The CIE76 color difference ({Delta}E) ranged from 36 to 112 and was indexed as the CIELAB distance from the black stripe pattern. Pupil size was continuously recorded and discomfort ratings were collected after each trial. Across the colored backgrounds, more uncomfortable stimuli evoked stronger pupil constriction, even though luminance was held nearly constant. As in our spatial-frequency study, this stimulus-level increase did not translate into stronger constriction among observers reporting higher overall discomfort: participants who rated the stimuli as more uncomfortable overall showed shallower constriction. This pattern was captured by the maximum-constriction response, which differentiated high-from low-discomfort observers and was significantly associated with individual discomfort ratings. Together with our previous findings, the results suggest that pupil responses scale along the tested stimulus axis, whereas individuals reporting greater visual discomfort exhibit less pronounced maximum constriction. HighlightsO_LIChromatic background modulated discomfort and pupil responses at similar luminance. C_LIO_LIAcross backgrounds, higher discomfort ratings tracked stronger pupil constriction. C_LIO_LIAcross observers, higher mean discomfort tracked weaker pupil constriction. C_LIO_LIThis two-level dissociation recurs across spatial and chromatic manipulations. C_LIO_LIPupillometry may complement subjective reports of visual discomfort. C_LI

18
A behaviourally normed database of 1,377 natural sounds for auditory cognition and neuroscience

Plegat, M.; Araujo Vitoria, M.; Marinato, G.; Tita, B.; van der Lans, C.; Pijfers, M.; Esposito, M.; Bertovic, M.-S.; Formisano, E.; Giordano, B. L.

2026-08-28 neuroscience 10.64898/2026.08.25.746933 medRxiv
Top 0.1%
5.3%
Show abstract

Natural-sound research requires stimulus sets that combine acoustic standardization with detailed behavioural characterization. We present 1,377 two-second sounds representing 240 expert-defined source--action classes. We call this database "MaMa Sounds", as it resulted from the collaborative effort of two academic teams in Maastricht and Marseille. The sounds were manually curated, segmented, sampled at 16 kHz, and labelled with a noun identifying the source and a verb identifying the action. We release deidentified trial-level identification and familiarity data together with multiple per-sound norms (e.g., identification accuracy, confidence and agreement; familiarity), along with overall norms derived with principal component analysis. Noun, verb, and joint noun--verb norms are provided as direct means and medians with the number of contributing observations. This battery preserves process-specific information, while two principal-component scores provide compact overall behavioural-identifiability measures derived from response ease, semantic correspondence, agreement, and familiarity. The repository also contains deterministic response-cleaning code, participant and reference Word2Vec representations, and code reproducing the public sound-level tables. The resource supports stimulus selection, matching, and continuous modelling in auditory cognition and neuroscience.

19
Age-corrected model for predicting pupil diameter in real-world conditions from melanopic equivalent daylight illuminance

Spitschan, M.

2026-08-11 neuroscience 10.64898/2026.08.05.742771 medRxiv
Top 0.1%
5.1%
Show abstract

PurposePupil diameter in daily life depends on both the light reaching the eye and the observers age, but established prediction formulas require laboratory quantities that are rarely measured in natural environments. We developed a compact age-corrected model that predicts pupil diameter from melanopic equivalent daylight illuminance (mEDI). MethodsWe used an existing field dataset in which binocular pupil diameter and near-corneal spectral irradiance were recorded while 83 adults aged 18-87 years moved through indoor and outdoor environments. The analysis included 10,082 valid paired observations. We fitted a bounded sigmoid relating pupil diameter to mEDI and age, with each participant given equal influence, and assessed prediction in participants excluded from model fitting. Performance was compared with simpler models, a flexible generalised additive model (GAM), and Watson-Yellott predictions based on assumed field geometry. ResultsPupil diameter decreased smoothly as mEDI increased. Age primarily reduced the difference between pupils in dim and bright conditions, by 0.768 mm per decade, while the predicted bright-light diameter changed little with age. In held-out participants, the bounded model had a participant-balanced root mean squared error (RMSE) of 0.630 mm and mean absolute error of 0.537 mm. The GAM had a slightly lower point-estimate RMSE of 0.610 mm, but the difference was small and uncertain. The bounded model outperformed the tested log-linear, reduced, age-only, and Watson-Yellott alternatives. ConclusionAge and mEDI are sufficient to provide useful population-average pupil predictions across the observed adult age and real-world light range. The model is transparent, physiologically bounded, and nearly as accurate as a flexible GAM, but predictions approaching darkness remain uncertain because valid mEDI measurements were not available in that range. Key pointsO_LIA compact equation predicts population-average pupil diameter from age and mEDI alone. C_LIO_LIAge mainly compresses the pupils response range by reducing pupil diameter under dimmer conditions. C_LIO_LIPrediction error in unseen participants was close to that of a flexible GAM, without requiring a fitted smooth object. C_LIO_LIThe model is intended for the observed adult age and field-light range, not for extrapolation into darkness. C_LI

20
Beyond Accuracy: Reliability-Aware Cross-Farm Evaluation of Dairy Cow Vocalization Models

Kate, M.; Neethirajan, S.

2026-06-22 animal behavior and cognition 10.64898/2026.06.17.732832 medRxiv
Top 0.1%
5.0%
Show abstract

Automated analysis of dairy cow vocalizations has largely relied on supervised classifiers evaluated within a single farm, a setting that inflates apparent performance and gives no measure of how far predictions can be trusted. We address this with a three-layer framework that separates acoustic structure discovery, proxy-state inference, and reliability assessment, evaluated on 569 annotated clips from three commercial dairy farms. A frozen self-supervised speech encoder, latent-space segmentation, and stability-guided clustering convert continuous recordings into discrete acoustic units without behavioral labels. Proxy-state signal is then tested under audio-only, audio-plus-context, and leave-one-farm-out (LOFO) protocols designed to separate transferable acoustic structure from farm-specific shortcuts. The results suggest that cross-farm generalizability differs substantially across biologically distinct vocalization categories. Non-vocal physiological sounds transfer across farms (LOFO macro-F1 = 0.763) and calibrate well (expected calibration error reduced from 0.087 to 0.023), whereas resource-related calls collapse to a majority-class baseline (macro-F1 = 0.500) and distress-related calls degrade under farm holdout. Selective prediction improves the retained-set score of the multiclass functional proxy (0.407 to 0.430), and an end-to-end convolutional baseline matches or exceeds the framework on raw accuracy for the easier targets yet yields a roughly two- to six-fold larger calibration error and offers no abstention. Random cross-validation consistently overstates cross-farm utility. These findings show that acoustic models for livestock monitoring require reliability-aware evaluation rather than flat classification.